The Journal of the Acoustical Society of America
● Acoustical Society of America (ASA)
Preprints posted in the last 90 days, ranked by how well they match The Journal of the Acoustical Society of America's content profile, based on 35 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Bozdogan, A.; Aarts, R. M.
Show abstract
Elephants and other large mammals produce low-frequency vocalizations extending well below the 20 Hz lower limit of human hearing, a regime known as infrasound. These rumbles serve vital social and reproductive functions over distances of several kilometers, yet they are inaudible to human observers and cannot be reproduced by conventional small loudspeakers. We present a complete signal-processing pipeline that renders sub-20 Hz elephant rumbles perceptible through a small loudspeaker by exploiting the missing-fundamental psychoacoustic effect. Butterworth bandpass filters isolate the infrasonic content; a full-wave integrator nonlinear device (NLD) generates the harmonic series required for virtual pitch perception; and a hysteresis-comparator fundamental-frequency estimator normalizes the NLD output. The pipeline was validated on African elephant field recordings and deployed on a credit-card-sized, low-cost single-board computer with an infrasound microphone and a small Bluetooth loudspeaker, demonstrating live operation in the field. The processed output shows a 10 dB to 15 dB elevation in the loudspeakers efficient band during call segments compared with background. The system enables zoo visitors and wildlife observers to perceive elephant rumbles in real time, opening new avenues for behavioral studies and public engagement with animal communication.
Delaram, V.; Ananthanarayana, R. M.; Trine, A.; Miller, M. K.; Stecker, G. C.; Buss, E.; Monson, B. B.
Show abstract
Several types of cues contribute to speech recognition in multi-talker environments. In this study, we investigated how talker head-orientation related (THOR) cues and extended high- frequency (EHF; >8kHz) cues affect speech-in-speech recognition for both female and male speech. We examined the THOR benefit associated with a non-facing masker talker head orientation (relative to a facing orientation) as a function of masker talker facing angle. The target talker always faced the listener, whereas co-located maskers were tested with eight different masker head angles, ranging from 0{degrees} (facing the listener) to facing 180{degrees} away. Two filtering conditions were tested: full- band and low-pass filtered at 8 kHz. A THOR benefit was observed at masker head angles greater than 45{degrees}, increasing from 2 dB to 8 dB between angles of 67.5{degrees} and 180{degrees}. This benefit was reduced for low-pass filtered speech. Access to EHF cues improved performance, but only for masker head angles >22.5{degrees}. There was no significant relationship between 16-kHz pure-tone thresholds and performance for young, normal-hearing listeners with good EHF hearing. These findings indicate that listeners benefit from non-facing masker talker head orientations >45{degrees} when the target talker is facing the listener, with greater benefit for larger head angles.
Benecke, J.; Whitmer, W. M.
Show abstract
In conventional hearing-aid personalisation, clinicians cannot hear what their patients hear, and patients cannot often reliably detect or describe what they hear. Self-adjustment avoids this issue but requires user controls that adjust hearing-aid signal processing parameters to be effective, efficient and easy. In this study, we explored (a) the roles of interface complexity and stimulus type in the self-adjustment of hearing-aid gain, and (b) how well individuals can adjust one sound to match another to assess the same interfaces and stimuli. Adult hearing-aid users with mild to moderate symmetrical sensorineural hearing loss repeatedly adjusted the gain (a) to their preference from individual prescription (n = 41) and (b) to match their previous preferences from a random starting point (n = 32) using three interfaces representing different bass/mid/treble configurations and three stimuli (music, speech and speech-in-noise). The large interindividual variability in self-adjusted gains clustered into three patterns of deviation from initial prescription: increased relative bass, overall gain reduction, and close to initial prescription. There were no substantial effects of interface nor stimulus on self-adjustment reliability (median {sigma} = 2.8 dB), whereas absolute sound-matching error increased with increasing interface complexity and centre frequency. Neither individual matching accuracy nor questionnaire responses predicted either self-adjusted gains or reliability. Overall, these results show that many - but not all - hearing-aid users can adjust gains with reasonable reliability, and while it can be difficult to predict the behaviour from the individual, the individual applies a similar self-adjustment behaviour across different interfaces and stimuli.
Parra Pena, J. A.; Sorolla, C.; Quinteros Veas, N. F.; Ibarra, E. J.; Alzamendi, G. A.; Peterson, S. D.; Weerathunge, H. R.; Guenther, F. H.; Zanartu, M.
Show abstract
Accurate modeling of laryngeal motor control is key to understanding typical and disordered voice production. However, traditional biomechanical plant models based on ordinary differential equations (ODEs) often involve high computational costs and numerical instabilities, limiting their use in real-time closed-loop control frameworks. This study evaluates feature-driven machine learning (ML) regressors, specifically Random Forest (RF), Multilayer Perceptron Neural Networks (NN), and Polynomial Regression (PR), as surrogate forward models mapping laryngeal motor inputs to fundamental frequency and sound pressure level. Training data were generated with two biomechanical vocal fold models: the extended body-cover and the triangular body-cover. Results demonstrate that ML surrogates reduce execution times from seconds to milliseconds (e.g., 2 ms for PR), enabling stable real-time tracking via inverse Jacobian control. While RF provides the highest accuracy, NN and PR offer smoother control signals and smaller memory footprints. A practical performance threshold was identified near N = 1,000 training samples, below which accuracy degraded substantially when models were trained from scratch. These findings support ML surrogates as efficient and adaptable alternatives to direct numerical simulation, providing a foundation for future subject-specific modeling through transfer learning in data-limited clinical scenarios.
Danner, T.; Vyshnevetska, V.; Friedrichs, D.; Moran, S.
Show abstract
Perceptual experiments show that listeners recognize female speakers with lower accuracy than male speakers. Automatic speaker recognition systems may also show performance bias against female speakers even when training data sets are gender balanced. The underlying reasons for this discrepancy are unclear. Here, we apply geometric morphometrics to quantify sex-related morphological vocal tract disparity -- the extent of shape variation -- across both resting and articulatory configurations. We find that male speakers exhibit greater disparity in both resting and articulatory configurations. This morphological idiosyncrasy may in turn generate more discriminable acoustic signatures and offer a biological explanation for higher recognition accuracies for male voices by humans and machines. Our results suggest that innate variation in vocal tract morphology may contribute to performance bias in voice technology and voice perception by human listeners.
Colak, H.; Guo, X.; Benzaquen, E.; Gurusiddappa, M.; Banerjee, A.; Choi, I.; Sedley, W.; Griffiths, T. D.
Show abstract
ObjectivesOutcomes following cochlear implantation vary substantially across adult recipients, and the cognitive and perceptual factors contributing to this variability are not fully understood. This poses a challenge for developing strategies to improve cochlear implant outcomes, as such approaches require a clearer understanding of the mechanisms underlying individual listening difficulties. In this study, we investigated auditory cognitive measures in cochlear implant (CI) users to further elucidate the origins of this variability. DesignThirty-seven adult cochlear implant users completed measures of auditory cognition, comprising auditory working memory (AWM) and sound segregation ability, measured using an auditory figure-ground task (AFG), as well as measures of peripheral temporal and spectral processing, comprising the temporal modulation detection threshold (TMDT) and spectral ripple discrimination threshold (SRDT). Speech perception outcomes were assessed using word-in-noise (WIN) and sentence-in-noise (SIN) tasks. Separate multiple linear regression models evaluated the unique contribution of the auditory cognition measures to WIN and SIN performance, after accounting for the peripheral measures. ResultsBoth regression models explained a substantial proportion of variance in speech-in-noise outcomes (WIN: adjusted R{superscript 2} = 0.55; SIN: adjusted R{superscript 2}=0.57, both p < 0.001). For WIN performance, AFG and AWM were significant predictors. A similar pattern was found for SIN performance, where lower AWM ability and poorer AFG segregation were linked to poorer sentence listening in noise. No significant effects of spectral ripple discrimination or temporal modulation detection were observed in either model, even though both were significantly correlated with WIN performance. ConclusionsThese findings indicate that auditory working memory and sound segregation ability are robust predictors of speech-in-noise outcomes in adult cochlear implant users, across both word- and sentence-level measures. Together, the results may help explain why speech-in-noise outcomes remain highly variable among CI users, even when basic sensory encoding abilities are taken into account. Incorporating measures of auditory working memory and fundamental sound segregation may therefore improve outcome prediction and help in developing more individualised rehabilitation strategies.
Mackey, C. A.; Mondul, J. A.; Ramachandran, R.
Show abstract
How sensory information is processed over time is often conceptualized as a process of temporal integration. Recently, auditory temporal integration has received renewed attention as a potential assay of hidden hearing loss caused by cochlear synaptopathy in rodent and avian studies. How these results relate to human hearing is in question due to a lack of studies in primates, and, more generally, the neural basis of auditory temporal integration is unclear, as most subcortical studies of it have been conducted under anesthesia. We have recently introduced a nonhuman primate (NHP) model which can address translational questions about auditory temporal integration and hidden hearing loss. Thus, in this study, we utilized single-unit recordings and compared derived neurometric measures to psychometric measures of temporal integration in normal hearing NHPs performing a tone-in-noise detection task. We then assessed psychometric measures of temporal integration in NHPs before and after noise exposure. In normal hearing NHPs, cochlear nucleus and inferior colliculus (IC) integration rates were significantly greater than psychometric rates. However, in noise only, [~]25% of IC neurons exhibited similar integration rates to behavior. After noise exposure, psychometric integration was disrupted for brief stimuli presented in quiet, but not in noise. The dynamic range of the psychometric function reliably increased, months after recovery from the noise-induced temporary threshold shift (TTS). Together, these data identify a subcortical neural substrate for temporal integration in noisy environments and suggest that behavioral assays of temporal integration may serve as sensitive indicators of subclinical hearing loss.
Azadpour, M.; Neukam, J.; Capach, N.; Svirsky, M.
Show abstract
Cochlear implants (CIs) restore hearing by stimulating auditory neurons to encode amplitude envelopes across frequency bands, providing essential cues for speech recognition. This study investigated how stimulation pulse rate constrains temporal envelope processing and speech cue perception in ten post-lingually deaf CI users by evaluating amplitude modulation (AM) detection thresholds and consonant identification performance across pulse rates. The effects of pulse rate on temporal processing and speech perception were examined using both standard clinical multi-channel strategies and single-channel strategies designed to isolate within-channel envelope representations. Results revealed a significant decline in AM detection and consonant recognition performance at the lowest tested pulse rate of 125 pulses per second (pps), consistent with perceptual constraints on temporal processing at low carrier rates, rather than inadequate envelope sampling. At the highest pulse rate of 4000pps, a non-significant reduction in AM detection was observed which may be consistent with previously reported reductions in amplitude discrimination at high pulse rates. Consonant recognition performance remained stable across clinically relevant pulse rates (250-2000pps), though listener-specific pulse rate effects were observed. Notably, significant correlations were found between single-channel and multi-channel performance in AM detection and consonant recognition tasks. These findings support an important contribution of within-electrode temporal envelope processing to multi-channel speech perception and highlight the clinical relevance of individual variability in pulse rate effects.
Dirks, C. E.; Guest, D. R.; Oxenham, A.
Show abstract
Context effects are ubiquitous across sensory systems and reflect a general encoding principle for both simple and complex stimuli. One simple context effect, contraction bias, manifests in two-interval perception tasks as a bias of the perceived magnitude of the first stimulus toward the center of the overall magnitude range. The underlying cause of contraction bias is unclear. One explanation is that a listeners magnitude estimate of the first stimulus is combined with a perceptual anchor, usually the mean stimulus magnitude, biasing it toward the anchor (sensory model). An alternative explanation is that a listeners response criterion shifts, based on the magnitude of the stimulus pair, relative to the mean magnitude of the stimuli range (decision model). Two pitch-discrimination experiments were performed to test these hypotheses in the auditory domain. The first was a forced-choice discrimination task, where listeners were asked to identify the higher or lower tone in a pair. The second was a same-different task where listeners indicated whether or not the two tones in a pair differed in frequency. Contraction bias was observed in the higher-lower discrimination task, even after extensive perceptual training with feedback. In contrast, no contraction bias was observed in the same-different task. Computational models of the sensory and decision hypotheses were fit to data from both experiments. The sensory model captured the pattern of results the higher-lower experiment but erroneously predicted a contraction bias in the same-different task. The decision model produced similar predictions to the sensory model in the higher-lower task but correctly predicted no contraction bias in the same-different task, and produced lower prediction errors and more stable parameter estimates in both paradigms. Overall, the results suggest that the underlying nature of the contraction bias may reflect decision, rather than sensory, biases based on the context.
Bilger, H.; J. Ryan, M.; Clarke, J.
Show abstract
The human larynx, compared to those of closely related primates, lies deeper in the throat and lacks vocal membranes and air sacs. These shifts are usually analyzed regarding their acoustic effects on vowel-like vocalizations, since the evolution of speech was long thought to require an expansion of vocal range driven by vocal tract modifications. However, vowels are just one type of phoneme, and speech is just one class of human utterance. To understand the evolutionary underpinnings of known shifts in human vocal morphology, a broader bioacoustic comparison is needed. Specifically, the range of sounds used in human speech must be compared to that employed in other human vocalizations and in the repertoires of extant close primate relatives. Here, we measure the acoustic-feature space occupied by human speech, non-linguistic, and musical vocalizations along with the calls of chimpanzees, bonobos, and chacma baboons. We use Mel-frequency cepstral coefficients to create an acoustic space depicting the spectro-temporal features of over 750,000 brief vocal segments sourced from published databases and other verified sources. Speech and song occupied significantly less volume in this acoustic space than human non-linguistic vocalizations. In addition, the acoustic-feature volumes of speech and song were not statistically distinct from those of non-human primates. These results suggest that speech was not enabled by an expansion of human vocal acoustic space. Anatomical shifts unique to humans may have led to an elaboration of non-linguistic utterances, but learned vocalizations use a surprisingly small fraction of this space. Our understanding of human vocal evolution will be further informed by additional systematic comparisons of the function and homology of non-speech vocalizations, along with the collection and incorporation of more complete non-human primate vocal datasets, especially from Gorilla and Orangutan.
Ghosh, A.; Borgohain, J.; War, R. M.; Rajaraman, B. K.
Show abstract
Ensiferans are nocturnal insects (Order Orthoptera) that produce mating advertisement calls using stridulatory organs on modified forewings. These calls, typically made by males, are species-specific and serve as indicators of forest health. In biodiverse ecosystems like the subtropical forests, caller density is high, and ecological constraints such as intra- and interspecific acoustic competition, masking interference, and predation pressure can influence calling behavior. These pressures lead to variation in call structures and differences in spatiotemporal acoustic space use, leading to variations in community call type composition across the seasons. Passive Acoustic Monitoring (PAM), a non-invasive and cost-effective technique, is widely used for long-term monitoring in vertebrate taxa, but is less commonly applied to terrestrial invertebrates. In this study, we employed PAM to quantify acoustic diversity, acoustic space use, separation of different call types, seasonal calling patterns, and seasonal variation in call type composition among nocturnal Ensiferan callers. Year-round recordings were conducted using AudioMoth devices in the Khasi Hills of Meghalaya, part of the Indo-Burma Biodiversity Hotspot, at a 48 kHz sampling rate. Acoustic samples were processed using Raven Pro software. We identified 33 distinct call types, mutually distinctly differing in spectral, temporal, or both parameters. Principal Component Analysis revealed fine-scale separation of call types. While there were seasonal shifts in call types, with the dry season having the least number of callers, overall call type composition remained stable across pre-monsoon and monsoon seasons. Acoustic Space Use (ASU) analysis indicated greater use of lower frequency bands consistent with ground cricket presence, as well as seasonal variation in spectral occupancy. This foundational study is the first of its kind in Northeast India and demonstrates the potential of PAM in studying invertebrate soundscapes.
Fish, E.; DiNino, M.
Show abstract
Acoustic cues such as pitch and spatial location allow listeners to attend to a target speaker and ignore competing talkers, aiding speech recognition in background noise. Diminished ability to utilize acoustic cues for speech stream segregation may thus contribute to older adults' challenges hearing in noise. Adults aged 18-74 completed a speech-in-speech identification task with three conditions containing 1) only pitch cues (fundamental frequency), 2) only spatial cues (interaural time differences; ITDs), and 3) both pitch and spatial cues for segregating a target talker from competing talkers. Hearing thresholds at standard and extended high frequencies (EHFs), auditory brainstem responses (ABRs), and digit span scores were acquired to examine the influence of sensory and cognitive factors on use of each acoustic cue for speech-in-speech recognition. Significant differences were observed between cue condition scores indicating that use of the available cue(s) drove performance. ABR metrics were not a significant predictor but digit span scores significantly predicted scores on all three cue conditions. Working memory abilities therefore set a baseline for participants' speech-in-speech recognition regardless of the acoustic content. Hearing thresholds at standard frequencies significantly predicted scores on the Pitch condition. EHF hearing thresholds better predicted Spatial and Both Cue condition performance, suggesting that EHF thresholds represent auditory processing important for coding ITDs. Age group analysis revealed that older adults (aged 40+) performed significantly more poorly on all cue conditions of the speech-in-speech recognition task relative to younger adults. Age-related changes in auditory sensory processing may therefore impair older adults' speech-in-noise perception by reducing their ability to use acoustic cues for segregating target and competing speech.
Marrone, J. P.; Ziliak, M. C.; Bartlett, E. L.
Show abstract
Auditory brainstem responses (ABRs) are a core part of objective functional evaluations of hearing sensitivity and subcortical auditory transmission. Manual assessments of ABR waveforms are still a primary means by which thresholds and peak amplitudes and latencies are measured, which is time-consuming and prone to user variability. Automated methods have offered promising alternatives for ABR classification, but they have sometimes been limited in accuracy or robustness. Here, we developed and tested a supervised convolutional neural network (CNN) based ABR peak classifier that works across sound levels and sound frequencies that can be run quickly on a personal computer using single or dual-channel ABR inputs. For ABR peaks I, III, IV, and V, the classifier achieved over 95% accuracy. High accuracy was maintained even after noise-exposure causing temporary or permanent threshold shifts, and over 90% of peaks were within 0.041 ms (1 sample) of the manually identified peak. Only a few hundred samples were needed to train the network, making it widely amenable to smaller data studies or where the number of subjects or sessions may be low.
Krasovskaya, S.; Coughlan, J. M.; Teng, S.
Show abstract
Some blind individuals use echolocation, a skill that allows them to better navigate their environment using echoes from self-generated mouth clicks reflected off surrounding surfaces. Echolocation involves a complex interplay of sensory accumulation, information processing, dynamic prediction, motor planning and execution in real-time. Computational modeling offers a valuable approach to understanding the cognitive and neural mechanisms underlying echolocation performance, in particular the temporal dynamics of the process. We present a computational model of human echolocation behavior based on a Kalman filter, where we treat the echolocator as an active sensor that maintains an internal belief about the target's location and continuously refines it via echo feedback. The model, based on observations of echolocation in blind human experts, simulates the use of mouth clicks and returning echoes to localize and orient toward a target under varying conditions. In the experiment, the target is placed at a random azimuth in the frontal plane. An echolocator aims a series of mouth clicks in various directions and infers the target azimuth using acoustic information received from the click echoes. The system integrates three major components: (1) a simulation of echoacoustic interaural time differences (ITD) to estimate the relative head-target angle; (2) a Kalman filter that processes these ITDs to iteratively update probabilistic beliefs about target location and associated uncertainty; and (3) a motor control system that modulates head movements with the current belief state. The Kalman filter serves as a representation of the internal state of the observer, where its beliefs drive the direction of head rotation, and its uncertainty estimates drive head velocity adjustments. Model performance demonstrates that simple predictive computational approaches can reproduce key aspects of echo-guided sensorimotor learning, providing a framework that may be leveraged to develop biologically plausible models, advance understanding of best practices, and potentially improve intervention strategies.
Plegat, M.; Araujo Vitoria, M.; Marinato, G.; Tita, B.; van der Lans, C.; Pijfers, M.; Esposito, M.; Bertovic, M.-S.; Formisano, E.; Giordano, B. L.
Show abstract
Natural-sound research requires stimulus sets that combine acoustic standardization with detailed behavioural characterization. We present 1,377 two-second sounds representing 240 expert-defined source--action classes. We call this database "MaMa Sounds", as it resulted from the collaborative effort of two academic teams in Maastricht and Marseille. The sounds were manually curated, segmented, sampled at 16 kHz, and labelled with a noun identifying the source and a verb identifying the action. We release deidentified trial-level identification and familiarity data together with multiple per-sound norms (e.g., identification accuracy, confidence and agreement; familiarity), along with overall norms derived with principal component analysis. Noun, verb, and joint noun--verb norms are provided as direct means and medians with the number of contributing observations. This battery preserves process-specific information, while two principal-component scores provide compact overall behavioural-identifiability measures derived from response ease, semantic correspondence, agreement, and familiarity. The repository also contains deterministic response-cleaning code, participant and reference Word2Vec representations, and code reproducing the public sound-level tables. The resource supports stimulus selection, matching, and continuous modelling in auditory cognition and neuroscience.
Wang, F.; Utianski, R. L.; Duffy, J. R.; Barnard, L. R.; Botha, H.
Show abstract
This study examined the extent to which goodness of pronunciation (GoP) scores and phonological posterior probabilities capture perceptual ratings of speech severity in individuals with motor speech disorders (MSD). Speech recordings of the word catastrophe were obtained from 489 participants, including 333 neurologically typical controls and 156 individuals with MSD. GoP scores were derived using traditional acoustic features and self-supervised speech representations, including WavLM and XLS-R, across multiple modeling approaches, while phonological posterior probabilities were extracted using Phonet. Model performance was evaluated using Kendall's rank correlations, regression, and receiver operating characteristic analyses against speech-language pathologists' perceptual ratings of sound distortion and intelligibility. Both GoP and phonological posterior probabilities were significantly associated with perceptual ratings. Self-supervised speech representations substantially outperformed traditional acoustic features, with WavLM-based GoP using k-nearest neighbors achieving the strongest performance. Across correlation, regression, and classification analyses, GoP consistently outperformed phonological posterior probabilities for both sound distortion and intelligibility. Age and gender had minimal influence on model-derived measures or their relationships with perceptual ratings. These findings demonstrate the value of self-supervised GoP as an objective measure of speech impairment while highlighting the complementary role of phonological posterior probabilities in characterizing articulatory aspects of motor speech disorders.
Dantanarayana, N. D.; Li, Y.; Litovsky, R. Y.; Borjigin, A.
Show abstract
Humans often communicate and learn in noisy, complex listening environments. Here, we investigated the effects of spatial hearing and semantic context cues on speech intelligibility and listening effort in young adults with typical hearing. The listening task included conditions in which target speech and speech maskers were either spatially co-located or separated. Target sentences were either semantically coherent or anomalous, while the masker comprised a mixture of two coherent sentences. Results showed higher speech intelligibility in spatially separated than co-located conditions, demonstrating a robust spatial release from masking (SRM), which is consistent with prior findings. SRM did not differ between semantically coherent and anomalous sentences, indicating comparable benefits of spatial cues across semantic contexts. However, within each spatial configuration, intelligibility was higher for coherent than anomalous sentences. Listening effort, indexed by peak pupil dilation in pupillometry measurement, was reduced in spatially separated conditions, suggesting a trend toward a release from listening effort. Analysis of the timing of peak pupil dilation revealed a significantly delayed peak dilation for anomalous sentences in the co-located condition compared with coherent sentences in the separated condition, indicating increased processing demands in the absence of spatial and semantic cues. Finally, SRM was correlated with the magnitude of release from listening effort for coherent sentences, but not for anomalous sentences, suggesting that intelligibility and listening effort benefits might co-occur when contextual cues are available.
Hauser, S. N.; Sivaprakasam, A. N.; Bharadwaj, H.; Heinz, M. G.
Show abstract
Purpose: Otoacoustic emissions (OAEs) are used to assess outer hair cell (OHC) function. Clinical interpretation of OAE responses, however, is often limited to a present/absent binary since both physiological factors and measurement variability affect the measured OAE amplitude. Prior work showed elevated OAE responses in sedated compared to awake chinchillas, pointing to the potential influence of the medial olivocochlear (MOC) efferents on amplitudes, but this finding is inconsistent across species and OAE type. Here, we aimed to further investigate the effect of anesthesia on distortion- and reflection-type emissions in chinchillas using swept stimuli and more reliable calibration methods. Methods: Swept distortion-product (DP) and stimulus-frequency (SF) OAEs were measured in chinchillas with and without ketamine/xylazine sedation. Stimuli were presented using in-ear forward pressure level calibrations. DPOAE and SFOAE amplitudes and estimated Qerb from SFOAE group delays were compared across the two conditions. Results: We found that low-frequency DPOAE amplitudes were elevated when animals were sedated. The difference in SFOAE amplitudes was more variable across animals but appeared mildly reduced in sedated animals. Qerb estimates were slightly higher in sedated animals at some frequencies. The effect of sedation was not different across sexes. Conclusion: Taken together, these findings suggest that sedation impacts OAE measurements in chinchillas. MOC modulation could account for the present findings and differences across species. For diagnostic precision, OAE responses should be considered in the context of not only intrinsic OHC function but also extrinsic physiological processes that can modulate OHCs.
Davies, T.; Bleeck, S.
Show abstract
Objective: This study investigated whether plosive consonants carry a perceptual loudness weighting that significantly exceeds that of non-plosive consonants when judged by hearing-impaired listeners. Design: A prospective loudness matching experiment utilizing the method of adjustment. Study Sample: 19 consenting native English speakers (Mean age: 61.4, SD: 16.4) with bilateral mild to moderate high-frequency sensorineural hearing loss, indicative of presbycusis. Stimuli: 13 vowel-consonant-vowel (VCV) nonsense syllables, exclusively utilizing the flanking vowel /u/. Results: Descriptive analysis revealed a strong time-order effect influencing loudness judgments for 7 of the 13 VCV test stimuli. Statistical testing showed no significant didference (P = 0.94) between the relative amplitudes corresponding to the point of equal loudness for plosive-containing versus non-plosive-containing VCV stimuli. However, 6 individual VCV stimuli, containing consonants from 4 separate manners of articulation, produced significant loudness matching data (P < 0.01). Conclusions: The results falsify the hypothesis that plosives, analyzed collectively as a class, possess a heavier perceptual loudness weighting than non-plosive consonants. While 6 individual VCV stimuli indicated potential individual consonantal loudness weightings, these findings must be interpreted cautiously due to the restriction to a single vowel context and the presence of procedural time-order biases.
Scott, M. T.; Limon, P. N.; Popelka, G. R.; Butts Pauly, K.; Norcia, A. M.; Ash, R. T.
Show abstract
Auditory confounds have proven to be a major hurdle in the elucidation of veridical neuromodulation effects with transcranial ultrasound stimulation (TUS). Auditory noise masks are an essential method to reduce the audibility of TUS and have shown promise in several studies. Here we describe a novel approach for design, calibration, and psychometric validation of auditory noise masks to reduce the perceptibility of TUS. White noise masks and spectrum-tuned masks matched to a TUS protocol that generates highly salient auditory costimulation (487.5 Hz pulse repetition frequency, 10% duty cycle, 68 W/cm2 pulse-peak average intensity, 500 kHz acoustic frequency) were generated, and dB(A) levels were calibrated with an artificial ear. The masker levels needed to reduce TUS detection performance in a two-interval forced choice task were determined with an adaptive QUEST+ staircase in 20 neurotypical participants. Detection performance of this highly salient TUS protocol was driven to near chance performance (<55%) in 17/20 participants with white-noise and 18/20 participants with spectrum-tuned noise. However, high masker levels approaching safety limits were needed to render TUS inaudible for the majority of participants, indicating the need for formal masker calibration for these TUS settings. Additionally, against expectation the spectrum-tuned masker did not significantly outperform the white-noise masker, suggesting that perceptibility of TUS auditory costimulation does not lawfully follow the sound expected from its pulse envelope and the known spectrum of human hearing.